Tag
1 article
This article explains how Perplexity's open-sourced Lily inference engine optimizes large language model execution on Apple Silicon using Rust and Metal, achieving significant performance gains over existing solutions.